JMIR Public Health and Surveillance
◐ JMIR Publications Inc.
Preprints posted in the last 7 days, ranked by how well they match JMIR Public Health and Surveillance's content profile, based on 45 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.
Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.
Show abstract
Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.
Davis, J. T.; Kaur, G.; Hines, A.; Ben-Nun, M.; Venkatramanan, S.; Brooks, L.; Mathis, S.; Ajelli, M.; Litvinova, M.; Kummer, A. G.; Ventura, P. C.; Mhade, S.; Weber, D.; Shemetov, D.; DeFries, N.; McDonald, D. J.; Yamana, T.; Zepeda-Tello, R.; Shaman, J.; Yaari, R.; Pei, S.; Webber, A.; Shandross, L.; Ray, E.; Wadsworth, S.; Niemi, J.; Redman, W. T.; Mullany, L.; Posner, R.; Mallela, A.; Lin, Y. T.; Hlavacek, W. S.; Smart, A.; Gill, A. A.; Drennan, A.; Fiebiger, B. J.; Miller, E. F.; Lee, J.; Mihaljevic, J. R.; Geist, K. A.; Baltz, M.; Bernik, O.; Truong, Y.-M. B.; Chen, Y.; Grosvenor, C. J.;
Show abstract
Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making.
Packard, S. E.; Russo, T.; Parrott, J.; Sisti, J.; Lans, A.
Show abstract
Objectives: To estimate the prevalence of Post-Exertional Malaise (PEM) among adults with prior COVID-19 and associated mental health and disability outcomes. Methods: We conducted a cross-sectional analysis of data from a survey of 9,620 adults with prior COVID-19 in New York City, collected May - June 2024. PEM was measured with the DePaul Symptom Questionnaire - Post Exertional Malaise, categorized by symptom duration (< 14 vs. [≥]14 hours). Weighted prevalence estimates were stratified by socio-demographic and clinical characteristics. Modified Poisson regression was used to assess the association of PEM with depression, anxiety, and disability. Results: The prevalence of PEM symptoms was 20.9% overall and 4.0% with symptom duration [≥]14 hours, representing over 800,000 New Yorkers affected and over 150,000 who meet a diagnostic criterion for ME/CFS. PEM prevalence was higher among women, transgender and non-binary adults, people of color, and lower educational attainment, chronic comorbidities, or disabilities. PEM was associated with 3 - 4 times higher prevalence of mental health outcomes and 4 - 5 times higher disability scores. Conclusions: PEM symptoms were common and strongly associated with disability and adverse mental health. Screening, pathways to care, and supportive policies are needed to mitigate long-term consequences, particularly among marginalized populations.
Amolo, P.; Mungai, L.; Karume, A. K.; Kibugi, J.; Mwende, W.; Botella, N.; Haldane, C.; Kamau, Y.; Marban-Castro, E.
Show abstract
Introduction Continuous Glucose Monitoring (CGM) is considered standard care in high-income countries. There is, however, limited published evidence on CGM use in low- and middle-income countries. The purpose of this study was to assess the usability, acceptability, and feasibility of CGM use among people living with type 1 diabetes (T1D) and caregivers in a low-resource setting. Research Design and Methods This prospective study conducted at the Kenyatta National Hospital purposively enrolled persons aged 4-25 years who had been on management for T1D for at least six months, and caregivers of those under 18 years. Fourty youth living with T1D used CGM for three months in place of self monitoring of blood glucose (SMBG). The System Usability Scale (SUS), a Theoretical Framework of Acceptability-based questionnaire, the Diabetes Distress Scale (DDS), the Glucose Monitoring Satisfaction Survey (GMSS), and a feasibility survey were administered. Outcomes were summarized descriptively, including means, medians, and frequencies using R statistical software. Results The median SUS score was 98.8 (IQR 92.5-100.0). Acceptability was high, and the median total GMSS score improved from 3.73 to 4.73. Among adolescents and adults, the median overall DDS score reduced from 1.54 to 1.36, with reductions in scores in all domains, except for hypoglycemia distress which increased, and physician distress which remained low. Among caregivers, the median overall DDS score declined from 2.05 (moderate distress) to 1.90 (low distress), with modest reductions in teen management and parent-teen relationship distress and a slight increase in personal distress. Median CGM active wear time was 89%. Conclusion This study comprehensively evaluated CGM across usability, acceptability, and feasibility outcomes, with the findings supporting the integration of CGM into routine diabetes management in low-resource settings. The short follow-up period, however, may not capture changing perceptions or long-term adherence.
Edmond, E. C.; Dreyer, A. J.; Winston, A.; Khoo, S. H.; Joska, J.; Nightingale, S.
Show abstract
Background Computerised cognitive testing may address the global challenge in identifying cognitive changes in people living with HIV scalably and affordably. We assessed a computerised battery (CB) of cognitive tests, in a prospective cohort (CONNECT) of people with HIV in a low-income peri-urban area of Cape Town, South Africa during a national programmatic switch from efavirenz- to dolutegravir-based antiretroviral therapy (ART). Methods We recruited 170 people with HIV and 91 people without HIV (controls) (140[82%] and 41[45%] followed up). The CB and gold-standard pen&paper cognitive testing (P&P) were performed at both timepoints. Technology familiarity/use questionnaire data were also collected. We compared performance in detecting lower group-level cognitive performance associated with efavirenz treatment. Furthermore, the CB was compared to P&P in classifying individuals with low cognitive performance, correlation of global test scores and domain-level scores between batteries, and practice effects between timepoints. Exploratory principal component analysis was also performed. Results People with HIV on efavirenz at baseline had lower performance on the computerised battery than controls, {Delta}T=2.6, p=0.0047. This difference was lost after switching to dolutegravir-based ART at follow-up. CB and P&P global T were moderately correlated (R2=0.203, p<0.001), and the CB performed moderately in classification of low cognitive performance against the gold standard (AUC 0.70, sensitivity 0.52, specificity 0.76, PPV 0.40, and NPV 0.84). Selecting the first three principal components improved both classification of low cognitive performance (AUC 0.77) and correlation strength with P&P global T (R2=0.3, p<0.001). The CB did not show practice effects. Most participants owned a mobile phone (95%, 85.9% of these smartphones). Performance was better in smartphone owners ({Delta}T=1.8) and computer owners (23%, {Delta}T=1.8). Conclusions Delivering computerised cognitive testing was feasible in this low-income southern African setting. The CB showed reasonable construct validity (detecting known lower cognitive performance associated with efavirenz-ART) and may detect broad cognitive characteristics such as processing speed and accuracy. However, correlation of CB results with gold standard P&P testing was low-moderate and may limit its applicability as a diagnostic tool. This might be improved by including a wider range of cognitive domains tested in the CB, or data driven analysis. Brief CBs may fulfil an initial screening role, followed by more detailed clinical assessment.
Pavia, M. J.; Amaro, I. F.; Xu, D.; Gonzalez-Hernandez, G.; Scotch, M.
Show abstract
Influenza vaccine effectiveness (VE) is estimated from a limited number of clinics using a test-negative design. These standard estimates face geographic, temporal, and operational constraints. Using Twitter/X data, we applied few-shot chain-of-thought prompting to identify self-reported vaccination status and influenza test results, then implemented a test-negative-like design to estimate VE. Our estimates fell within the range of interim reports and could complement current systems, improving feasibility, timeliness, and scalability.
Yendewa, G.; Chengsupanimit, T.; Dehghani, A.; Ahmed, A.; Mohareb, A.; Cohen, C.; Freeman, M.; Kim, H. N.; Ofotokun, I.; Dube, K.
Show abstract
Background: HIV/HBV coinfection is associated with substantial liver-related morbidity and mortality, yet the impact of social vulnerability (SV) on clinical outcomes has not been systematically assessed. We evaluated associations of multidimensional SV with mortality, hepatic, virologic, and extrahepatic organ outcomes among adults with HIV/HBV. Methods: We conducted a retrospective cohort study using TriNetX data from 110 U.S. healthcare organizations (2010-2026). We propensity score matched adults with HIV/HBV with and without documented SV 1:1 (2,024 per group). SV was defined using a four-domain framework encompassing material, healthcare access and engagement, interpersonal, and psychosocial vulnerability. Results: Over 15,900 person-years, SV was associated with higher mortality (hazard ratio [HR], 2.06; 95% confidence interval [CI], 1.72-2.47), liver composite events (HR, 1.37; 95% CI, 1.07-1.76), hepatic decompensation (HR, 1.94; 95% CI, 1.39-2.70), hepatic failure (HR, 2.39; 95% CI, 1.53-3.73), HBV viremia (HR, 1.69; 95% CI, 1.32-2.16), and HIV viremia (HR, 2.05; 95% CI, 1.71-2.46). SV was also associated with major adverse cardiovascular events (HR, 1.47), chronic kidney disease (HR, 1.49), and diabetes (HR, 1.25). Multidomain SV generally showed stronger associations than single-domain SV for most hepatic and virologic outcomes, with HR ranges of 1.76-2.62 versus 1.35-1.76 for single-domain SV. Healthcare access and engagement vulnerability was most consistently associated with mortality and hepatic outcomes. Conclusions: SV was associated with mortality, hepatic disease, impaired HIV/HBV control, extrahepatic organ morbidity, and acute care utilization in adults with HIV/HBV. SV assessment may improve risk stratification and identify actionable intervention targets during HIV/HBV care.
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
MURHABAZI BASHOMBWA, A.; TCHIO-NIGHIE, K. H.; NANA DJAPOU, M. C.; BUH NKUM, C.; BLAMA ABBA, I.; BEKOLO, C. E.; ATEUDJIEU, J.
Show abstract
Health facilities (HFs) routinely administer medicines and are expected to ensure patient safety by detecting, reporting, investigating, and analysing adverse events following exposure to drugs (AEFED). This study aimed to assess the implementation of pharmacovigilance activities in referral and regional health facilities in Cameroon and to identify pharmacovigilance training needs among healthcare personnel (HP). This was a cross-sectional descriptive study targeting referral and regional health facilities and healthcare personnel involved in patient care and pharmacovigilance activities in Cameroon. Health facilities were selected using stratified purposive sampling, while healthcare personnel were selected through exhaustive sampling. Data were collected using semi-structured electronic questionnaires administered face-to-face by trained enumerators. The questionnaires assessed the organization, resources, and implementation of pharmacovigilance activities at health facilities, as well as healthcare personnel knowledge of pharmacovigilance concepts, previous training, and perceived training needs. Of the 14 eligible health facilities, 10 (71.4%) consented to participate in the study. Of the 10 health facilities, 4 (40.0%) had an established pharmacovigilance unit, while 3 (30.0%) reported conducting neither detection nor notification activities. Among the 261 healthcare personnel approached, 214 (81.9%) participated. Only 41.6% had needed knowledge to detect an adverse event, while 72.9% were aware of adverse event notification procedures. Previous exposure to pharmacovigilance training was reported by 37.9% of healthcare personnel, and all participants expressed a need for additional training, particularly on national pharmacovigilance regulations (69.2%), organization of the pharmacovigilance system (67.3%), and adverse event detection (67.3%). The main reported challenges by healthcare personnel in the implementation of pharmacovigilance activities included insufficient budget allocation, limited access to pharmacovigilance training, lack of pharmacovigilance guidelines and insufficient qualified human resources. Pharmacovigilance implementation in referral and regional health facilities in Cameroon remains limited, with gaps in organizational structures, resources, healthcare personnel knowledge, and training. Strengthening pharmacovigilance systems through improved facility capacity, availability of essential tools, and targeted healthcare personnel training is needed to enhance drug safety surveillance.
Garcia Campos, M. A.; Rocha, T. A. H.; Perez de Souza, J. V.; Murase, L. S.; Murta, F.; Sartim, M. A.; Sachett, J.; Seabra de Farias, A.; Azevedo Machado, V.; Wen, F. H.; Staton, C. A.; Monteiro, W. M.; Gerardo, C. J.; Nickenig Vissoci, J. R.
Show abstract
Background: Snakebite envenoming is a major cause of preventable death and disability in the Brazilian Amazon, where long distances, sparse roads, and dependence on river transport delay access to antivenom. We developed location-allocation models to identify community health centers that could strategically expand access to antivenom in Amazonas State, Brazil. Methodology/Principal Findings: We conducted an ecological geospatial study using a 2025 WorldPop population surface, locations of existing and candidate health facilities, and a multimodal road-and-river transportation network derived from OpenStreetMap and HydroSHEDS. Population demand was represented by 7,065 populated centroids, including 1,586 within Indigenous territories. We applied a maximize-coverage algorithm with a six-hour travel-time threshold. Two models were developed: one for Amazonas excluding Manaus and one for populations living in Indigenous territories. Both models began with 77 facilities already providing antivenom and progressively added candidate community health centers until coverage gains plateaued. The plateau occurred at 110 facilities, corresponding to 33 additional centers. In the model excluding Manaus, this configuration covered 1,118,831 people, or 75.11% of the target population; 87.61% of those covered could reach care within three hours. In Indigenous territories, coverage increased from 50.55% to 69.50%, reaching 50,434 people, of whom 81.39% were within three hours of care. Validation used 3,595 snakebite notifications from the 30 highest-burden municipalities in the Brazilian Notifiable Diseases Information System during 2023-2025. The median proportion reaching care within six hours was 40.81% in observed data and 72.17% in model estimates. Conclusions/Significance: Strategically equipping 33 additional existing community health centers could substantially expand timely access to antivenom, particularly in rural and Indigenous areas. Location-allocation modeling that incorporates river transportation can support evidence-based decentralization of time-sensitive health services in geographically complex settings.
Ali, S. I.; Varatharajan, V.; Chacko, S. T.; Hazari, A.; Varghese, S. M.
Show abstract
Objectives This study aimed to assess sleep patterns and life satisfaction among employees of a private company in Dubai, United Arab Emirates, and to examine the relationships among sleep quality, life satisfaction, and selected demographic variables. A quantitative descriptive cross-sectional survey design was adopted. Methods A convenience sample of 110 male employees participated in the study. Data were collected using the Sleep Disorder Assessment Scale (16 items; Cronbachs = 0.89) and the Life Satisfaction Scale (5 items). Statistical analysis was performed using SPSS version 29, including descriptive statistics, chi-square tests, and Pearson correlation analysis. Results Most participants (66.4%) were aged 20-30 years, and 82.7% experienced moderate sleep-related problems. Mobile phone use before bedtime was common, with 60.9% reporting occasional use and 35.5% reporting regular use. Overall, 41.8% reported neutral life satisfaction, while 25.5% and 24.6% were slightly and extremely satisfied, respectively. A significant negative correlation was found between poor sleep patterns and life satisfaction (r = -0.389, p < 0.001). Mobile phone use before bedtime and shift work were significantly associated with sleep patterns (p = 0.048). Conclusion Poor sleep quality, particularly among shift workers and frequent bedtime mobile phone users, is associated with lower life satisfaction. Workplace interventions promoting sleep hygiene may enhance employee well-being.
Yao, R.; Wi, C.-I.; Beenken, M. J.; Watson, D.; Wheeler, P. H.; Finch, M.; Kelleher, D. P.; Anil, G.; Anderson, T.; Madden, K.; Okuno, S. H.; Odedina, F. T.; Westfall, E. C.; Park, E. Y.; Sharma, P.; Dugani, S.; Foss, R. M.; Hidaka, B. H.; Sosso, J. L.; Sabarish, S.; Singh, G.; Lugo-Fagundo, N.; Howick, J.; Kim, W. R.; Calvin, A. D.; Walker-Mcgill, C. L.; Rennert, L.; Juhn, Y. J.; Cerhan, J. R.; Lynch, B. A.
Show abstract
Purpose: This study assesses the association between colorectal cancer (CRC) screening and a validated, housing-based measure of individual-level socioeconomic status (SES, called HOUSES hereafter) within rural communities and determines whether HOUSES-integrated geospatial analysis can be used to tailor interventions. Methods: We used CRC screening data from a subset of Mayo Clinic Midwest patients living in cities without ready access to routine care in the Mayo Clinic Health System in 2019 to represent rural communities. At the individual level, we assessed the association between CRC screening rates and the HOUSES index, adjusting for age, sex, race/ethnicity, comorbidity, distance from home address to clinic, and area deprivation index, using a multilevel mixed-effects logistic regression model. Additionally, we conducted geospatial analysis to examine the correlation between hotspots of 1) lower CRC screening rates and 2) lower SES of the subject population (HOUSES quartile 1). Findings: Among 34,489 individuals (median age 64.0 years, 52.4% female), those with the lowest SES (HOUSES Q1) had 37% lower odds of being CRC screening adherent than those with the highest SES (HOUSES Q4) (adj. OR [95% CI]: 0.63 [0.58-0.69]). In the 14 identified HOUSES Q1 hotspots, there was a significant correlation in counts of HOUSES Q1 and low CRC screening (correlation coefficient=0.81). Conclusion: Lower SES was significantly associated with lower CRC screening among rural populations. HOUSES-enabled geospatial analysis identified geographic hotspots with lower CRC screening rates for targeted interventions to address disparities in CRC screening in rural communities. HOUSES may be a useful digital tool for cancer preventive care and research.
Marban-Castro, E.; Muhwava, L.; Girdwood, S.; Kemp, T.; Freitas, J.; Kamau, Y.; Otieno, M.; Akach, D.; Morato, A.; Sanz, S.; Fiechter, V.; Erkosar, B.; Watson, M.; Vetter, B.; Haldane, C.; Shilton, S.; Rheeder, P.; Dave, J. A.; Carrihill, M.; Karsas, M.
Show abstract
Introduction: Continuous glucose monitoring (CGM) offers an advancement over traditional self-monitoring of blood glucose (SMBG) for people living with type 1 diabetes (T1D). However, evidence on the acceptability and feasibility of different CGM use cases in African populations remains limited. Methods: This was a pragmatic three-arm, randomised controlled trial on CGM conducted among people living with T1D in three public healthcare clinics in South Africa. Participants were assigned to Arm 1 (continuous CGM), Arm 2 (periodic CGM), or Arm 3 (SMBG). Diabetes education was provided at all study visits. Feasibility was assessed by adherence to CGM use and through the Glucose Monitoring Satisfaction Survey (GMSS). Diabetes distress was measured by the Diabetes Distress Scale (DDS), health-related quality of life (HRQoL) by the EQ-5D scales, and acceptability using the Theoretical Framework of Acceptability (TFA). Surveys were collected on paper and transferred to OpenClinica. Analyses were performed in R. The trial was registered in the Clinical Trials Registry (NCT05944718) on July 13, 2023. Results: A total of 83 participants were included in Arm 1, 85 in Arm 2, and 80 in Arm 3. CGM mean active time was 55% in Arm 1 versus 69% in Arm 2. The proportion of participants meeting the [≥]70% active time threshold was higher in Arm 2 (52%) than in Arm 1 (34%). Diabetes' distress declined across arms during the intervention period, with no significant difference between arms; distress increased slightly six months post-intervention but remained below baseline. At 6 months, glucose monitoring satisfaction was significantly higher in both CGM arms than in the SMBG arm, and satisfaction increased over time in CGM arms. Health-related quality of life remained stable across arms during the intervention period with no significant difference between arms. High acceptability was observed in both CGM arms, with higher ratings in the periodic arm. Conclusions: CGM was acceptable to people living with type 1 diabetes and feasible to use in public-sector clinics in South Africa, with high acceptability under continuous and periodic use. Health-related quality of life remained stable across arms, and diabetes-related distress declined, during the intervention period, across arms. Glucose monitoring satisfaction rose significantly in both CGM arms compared to SMBG. Periodic CGM might be a promising and potentially more scalable option than continuous use for public-sector care.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Chimpandule, T.; Tweya, H.; Goeke, L.; Masina, T.; Macheso, S.; Low, N.; Jahn, A.; Imai-Eaton, J. W. W.
Show abstract
Background: In 2019, WHO recommended three consecutive reactive serological test results for HIV diagnosis to reduce false-positive diagnoses. Malawi changed from a two-test to a three-test strategy in 2022 as HIV test positivity declined. We assessed diagnostic performance, implementation fidelity, and costs. Methods: We analysed national HIV testing data from Nov 1, 2022, to Oct 31, 2025. Using observed three-test classifications as the reference standard, we reconstructed classifications under the two-test strategy. We estimated positive predictive value (PPV), implementation fidelity, potential false-positive diagnoses prevented, incremental costs, and time to offset testing costs through avoided antiretroviral therapy expenditure. Results: Among 9,885,599 encounters eligible for implementation-fidelity analysis, 99.98% followed a valid three-test pathway. The diagnostic-performance analysis included 9,862,908 encounters, of which 171,351 (1.7%) were classified HIV-positive and 9,138 (0.09%) were inconclusive. Under the two-test strategy, 1,209 inconclusive encounters with a T1+/T2+/T3- sequence would have been classified as HIV-positive. Retesting and reference-laboratory data indicated that 82.5% of these would subsequently be classified as HIV-negative, corresponding to 997 false-positive diagnoses prevented (10.3 per 100 000 three-test non-positive encounters; 95% CI 9.7-10.9). Retesting within 1-2 weeks was associated with the highest odds of potential false-positive classification (adjusted OR 39.37, 95% CrI 30.63-50.61). The incremental cost was US$471 per false-positive diagnosis averted and was offset within 7.30 years. Conclusions: Malawi's transition to a three-test HIV testing strategy prevented false-positive diagnoses and unnecessary antiretroviral therapy at modest cost, supporting broader adoption of WHO guidance in similar settings. Funding: Gates Foundation.
Catrianiningsih, D.; Felisia, F.; Abdalla, A. S.; Puspitasari, S.; Dwihardiani, B.; Mulia, H. N.; Hidayat, A.; Triasih, R.
Show abstract
In primary healthcare centers lacking advanced imaging, community-based active tuberculosis (TB) case finding often relies on basic symptom screening. This approach often misses cases and leads to the inefficient allocation of rapid molecular testing (RMT). We aimed to develop and internally validate a simple clinical triage scorecard to improve TB detection and guide RMT use in resource-constrained settings. We conducted a retrospective cross-sectional study of 15,137 adults ([≥]18 years) evaluated within the Zero TB Yogyakarta program (2020-2025). Participants with complete clinical assessments and confirmatory GeneXpert results were included. Using multivariable logistic regression, we identified independent clinical predictors, which were subsequently transformed into an integer-based point scorecard. Model performance was evaluated via discrimination and calibration, utilizing bootstrap resampling (1,000 iterations) for internal validation. Among the 15,137 participants, 251 (1.7%) were GeneXpert-positive. The final multivariable model identified eight independent predictors: age, male sex, body mass index, prolonged cough, hemoptysis, unexplained weight loss, TB contact history, and diabetes mellitus. The model demonstrated strong predictive accuracy, with an optimism-adjusted AUROC of 0.836 and good calibration. When translated to the integer scorecard and compared directly to standard national symptom screening, the scorecard performed (AUROC 0.81 vs. 0.73; p<0.001). At a high sensitivity cut off score of [≥] 0, the tool achieved 93.63% sensitivity and 41.33% specificity. This point-of-care clinical scorecard provides higher diagnostic accuracy than standard symptom screening algorithms. By offering flexible operational thresholds, it empowers local health programs to dynamically balance the urgency of case detection with available diagnostic capacity, optimizing GeneXpert allocation where advanced radiological imaging is unavailable.
Oshinubi, K.; Covington, J.; Busser, N.; Townsend, J.; Will, J.; Ruberto, I.; Kretschmer, M.; Chen, Y.; Doerry, E.; Hepp, C. M.; Mihaljevic, J. R.
Show abstract
Mosquito-borne diseases pose a growing public health challenge as climate change reshapes vector population dynamics. West Nile virus (WNV), transmitted between birds and Culex mosquitoes, disproportionately affects Maricopa County, Arizona, one of the nation's highest-burden counties, yet whether models that include weather and avian dynamics improve forecast accuracy remains unclear. Using a 15-year weekly time series of mosquito abundance, mosquito infection prevalence, and human cases, we developed four mechanistic model configurations of varying complexity, from mosquito-human dynamics alone to full models incorporating avian dynamics and weather forcing. We fitted each model to the weekly-observed data, generated probabilistic 1- and 2-week-ahead forecast horizons, and evaluated forecasts against a historical baseline. All configurations fit the data equally regardless of weather or avian dynamics. However, models incorporating both birds and weather created more accurate forecasts of mosquito abundance and mosquito infection prevalence, and all configurations outperformed the baseline for forecasting human cases. Forecast accuracy was highest in summer and fall, and ensemble aggregation sometimes outperformed every individual model, stabilizing predictions across the 15-year record. These findings indicate that avian and weather dynamics are most critical for predicting mosquito-specific data, positioning this framework as a scalable tool for public health planning for WNV surveillance under climate change.
Hessel, M.; Inda Diaz, J. S.; Sjöberg, A.; Salva-Serra, F.; Helldal, L.; Jirstrand, M.; Johnning, A.; Kristiansson, E.; Skovbjerg, S.
Show abstract
Antimicrobial resistance is a public health challenge, driving the need for rapid, cost-effective diagnostic support tools. Artificial intelligence (AI) may enable prediction of susceptibility to untested antibiotics from known susceptibility results, but prospective clinical validation is required before routine use. We evaluated an AI-based decision support method, trained on invasive isolates from the European Surveillance System (TESSy), for prediction of antibiotic susceptibility in clinical Escherichia coli urine isolates. The evaluation included 99 E. coli isolates from urine samples with diversity in age, sex, and antibiotic susceptibility. Predictions were evaluated for 14 antibiotics using patient metadata and susceptibility results for 4-8 antibiotics as input. Prediction uncertainty was handled using conformal prediction, allowing abstention when confidence was insufficient. EUCAST disk diffusion test results were used as reference and genomic sequence data was used to explore mechanisms of the AI performance. Without conformal prediction, 84% of predictions were correct when susceptibility results of six antibiotics were used to predict susceptibility to eight additional antibiotics. Across all predictions generated using susceptibility results for six antibiotics as input, the major and very major error rates were 19% and 12%, respectively. Prediction errors varied between antibiotics and were associated with certain phenotypic and genotypic resistance patterns. Conformal prediction reduced errors but increased abstentions; at confidence levels of 90%, 95%, and 97.5%, the model abstained in 9.6%, 14%, and 22% of instances. The method showed promising performance, but its clinical use remains limited and may require diagnostic data beyond susceptibility test results and demographic variables.
Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.
Show abstract
BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.